AITopics | learning compositional neural program

Learning Compositional Neural Programs with Recursive Tree Search and Planning

Neural Information Processing SystemsDec-25-2025, 18:07:48 GMT

We propose a novel reinforcement learning algorithm, AlphaNPI, that incorporates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and increase interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning. The experiments show that AlphaNPI can sort as well as previous strongly supervised NPI variants. The AlphaNPI agent is also trained on a Tower of Hanoi puzzle with two disks and is shown to generalize to puzzles with an arbitrary number of disks. The experiments also show that when deploying our neural network policies, it is advantageous to do planning with guided Monte Carlo tree search.

learning compositional neural program, recursive tree search, recursive tree search and planning, (7 more...)

Neural Information Processing Systems

Country: Asia > Vietnam > Hanoi > Hanoi (0.27)

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.83)
Information Technology > Artificial Intelligence > Representation & Reasoning > Search (0.63)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.51)

Add feedback

Learning Compositional Neural Programs with Recursive Tree Search and Planning

Neural Information Processing SystemsMay-27-2025, 13:37:00 GMT

We propose a novel reinforcement learning algorithm, AlphaNPI, that incorpo- rates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and in- crease interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning.

learning compositional neural program, recursive tree search and planning, specification, (4 more...)

Neural Information Processing Systems

Country: Asia > Vietnam > Hanoi > Hanoi (0.09)

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.88)
Information Technology > Artificial Intelligence > Representation & Reasoning > Search (0.79)

Add feedback

Reviews: Learning Compositional Neural Programs with Recursive Tree Search and Planning

Neural Information Processing SystemsJan-25-2025, 20:19:51 GMT

It instead learns the hierarchy of program subroutines in a curriculum fashion, adding a pre- and post-condition to each subroutine and extending the MCTS setup of AlphaZero to handle recursive subroutine calls. The paper demonstrates that the resulting formulation learns the programs in both Sorting and TowersOfHanoi domains more effectively than prior work.

learning compositional neural program, strong supervision, supervision, (9 more...)

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Representation & Reasoning > Search (0.40)

Add feedback

Reviews: Learning Compositional Neural Programs with Recursive Tree Search and Planning

Neural Information Processing SystemsJan-25-2025, 20:19:41 GMT

The authors should be commended for an excellent submission to NeurIPS. The concerns about clarity the reviewers raised seem to be addressable as the authors describe in their rebuttal. The topic: "unsupervised" (really, less-supervised) structured neural program induction is perfect for NeurIPS and the empirical results on sorting and other tasks as compared to the original neural programmer interpreter are exciting.

learning compositional neural program, recursive tree search and planning

Neural Information Processing Systems

Technology: Information Technology > Artificial Intelligence > Representation & Reasoning > Search (0.40)

Add feedback

Learning Compositional Neural Programs with Recursive Tree Search and Planning

Neural Information Processing SystemsOct-10-2024, 13:20:56 GMT

We propose a novel reinforcement learning algorithm, AlphaNPI, that incorpo- rates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and in- crease interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning.

learning compositional neural program, recursive tree search and planning, specification, (4 more...)

Neural Information Processing Systems

Country: Asia > Vietnam > Hanoi > Hanoi (0.09)

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.88)
Information Technology > Artificial Intelligence > Representation & Reasoning > Search (0.79)

Add feedback

Learning Compositional Neural Programs for Continuous Control

#artificialintelligenceSep-15-2020, 00:50:13 GMT

We propose a novel solution to challenging sparse-reward, continuous control problems that require hierarchical planning at multiple levels of abstraction. Our solution, dubbed AlphaNPI-X, involves three separate stages of learning. First, we use off-policy reinforcement learning algorithms with experience replay to learn a set of atomic goal-conditioned policies, which can be easily repurposed for many tasks. Second, we learn self-models describing the effect of the atomic policies on the environment. Third, the self-models are harnessed to learn recursive compositional programs with multiple levels of abstraction. The key insight is that the self-models enable planning by imagination, obviating the need for interaction with the world when learning higher-level compositional programs. To accomplish the third stage of learning, we extend the AlphaNPI algorithm, which applies AlphaZero to learn recursive neural programmer-interpreters. We empirically show that AlphaNPI-X can effectively learn to tackle challenging sparse manipulation tasks, such as stacking multiple blocks, where powerful model-free baselines fail.

large language model, machine learning, reinforcement learning, (10 more...)

#artificialintelligence

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.66)
Information Technology > Artificial Intelligence > Natural Language > Large Language Model (0.40)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.40)

Add feedback

Learning Compositional Neural Programs with Recursive Tree Search and Planning

PIERROT, Thomas, Ligner, Guillaume, Reed, Scott E., Sigaud, Olivier, Perrin, Nicolas, Laterre, Alexandre, Kas, David, Beguir, Karim, Freitas, Nando de

Neural Information Processing SystemsMar-19-2020, 02:46:10 GMT

We propose a novel reinforcement learning algorithm, AlphaNPI, that incorpo- rates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and in- crease interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning.

learning compositional neural program, recursive tree search and planning, specification, (4 more...)

Neural Information Processing Systems

Country: Asia > Vietnam > Hanoi > Hanoi (0.09)

Technology:

Information Technology > Artificial Intelligence > Machine Learning > Reinforcement Learning (0.89)
Information Technology > Artificial Intelligence > Representation & Reasoning > Search (0.80)

Add feedback

Filters

Collaborating Authors

learning compositional neural program

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

Learning Compositional Neural Programs with Recursive Tree Search and Planning

Learning Compositional Neural Programs with Recursive Tree Search and Planning

Reviews: Learning Compositional Neural Programs with Recursive Tree Search and Planning

Reviews: Learning Compositional Neural Programs with Recursive Tree Search and Planning

Learning Compositional Neural Programs with Recursive Tree Search and Planning

Learning Compositional Neural Programs for Continuous Control

Learning Compositional Neural Programs with Recursive Tree Search and Planning